Improving Centroid-based Text Classification Using Term-distribution-based Weighting System and Clustering
نویسندگان
چکیده
Centroid-based text classification is one of the most popular supervised approaches to classify texts into a set of pre-defined classes with relatively low computation. Based on the vector-space model, the performance of this classification particularly depends on the way to weigh terms in documents in order to construct a representative class vector for each class and degree of spherical shape in class. In this paper, we propose a method to improve classification accuracy by considering a number of statistical term weighting systems based on term-distribution, including factors of intra-class, inter-class, overall term frequency distribution and term-length normalization. An improvement using a clustering technique called hierarchical EM is also investigated. A number of experiments using drug information web pages and newsgroups data set, are made. The results show that our method outperforms standard tf-idf centroid-based, k-nearest neighbor and naïve Bayesian classifiers to some extent.
منابع مشابه
Using Class Frequency for Improving Centroid-based Text Classification
Most previous works on text classification, represented importance of terms by term occurrence frequency (tf) and inverse document frequency (idf). This paper presents the ways to apply class frequency in centroid-based text categorization. Three approaches are taken into account. The first one is to explore the effectiveness of inverse class frequency on the popular term weighting, i.e., TFIDF...
متن کاملEmpirical Evaluation of Centroid-based Models for Single-label Text Categorization
Centroid-based models have been used in Text Categorization because, despite their computational simplicity, they show a robust behavior and good performance. In this paper we experimentally evaluate several centroidbased models on single-label text categorization tasks. We also analyze document length normalization and two different term weighting schemes. We show that: (1) Document length nor...
متن کاملCombining homogeneous classifiers for centroid-based text classification
Centroid-based text classification is one of the most popular supervised approaches to classify texts into a set of pre-defined classes. Based on the vector-space model, the performance of this classification particularly depends on the way to weight and select important terms in documents for constructing a prototype class vector for each class. In the past, it was shown that term weighting us...
متن کاملA Joint Semantic Vector Representation Model for Text Clustering and Classification
Text clustering and classification are two main tasks of text mining. Feature selection plays the key role in the quality of the clustering and classification results. Although word-based features such as term frequency-inverse document frequency (TF-IDF) vectors have been widely used in different applications, their shortcoming in capturing semantic concepts of text motivated researches to use...
متن کاملImproving the Operation of Text Categorization Systems with Selecting Proper Features Based on PSO-LA
With the explosive growth in amount of information, it is highly required to utilize tools and methods in order to search, filter and manage resources. One of the major problems in text classification relates to the high dimensional feature spaces. Therefore, the main goal of text classification is to reduce the dimensionality of features space. There are many feature selection methods. However...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
دوره شماره
صفحات -
تاریخ انتشار 2001